Papers with zero-shot TTS models
Zero-Shot Text-to-Speech for Vietnamese (2025.acl-short)
Copied to clipboard
| Challenge: | Text-to-speech (TTS) synthesis has seen significant advancements in recent years. |
| Approach: | They propose to use PhoAudiobook to curated 941 hours of high-quality audio for Vietnamese text-to-speech models. |
| Outcome: | The proposed model improves on VALL-E, VoiceCraft, and XTTS-V2 models, highlighting their robustness in handling diverse linguistic contexts. |
Cross-Domain Audio Deepfake Detection: Dataset and Analysis (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing audio deepfake detection datasets are outdated and lack generalization capabilities. |
| Approach: | They construct a new cross-domain audio deepfake detection dataset comprising over 300 hours of speech data that is generated by five advanced zero-shot TTS models. |
| Outcome: | The proposed models achieve 4.1% and 6.5% error rates in the cross-domain ADD dataset generated by five advanced zero-shot TTS models. |